Papers with unimodal baseline models
What do Models Learn From Training on More Than Text? Measuring Visual Commonsense Knowledge (2022.acl-srw)
Copied to clipboard
| Challenge: | Existing evaluation methods to measure what language models learn from multimodal training are lacking. |
| Approach: | They propose two evaluation tasks to measure commonsense knowledge in language models by using visual data to evaluate multimodal models and unimodal baselines. |
| Outcome: | The proposed evaluation tasks show that training on a visual modality improves on the visual commonsense knowledge in language models. |